NSF PAR Search | NSF Public Access Repository

Note: When clicking on a Digital Object Identifier (DOI) number, you will be taken to an external site maintained by the publisher. Some full text articles may not yet be available without a charge during the embargo (administrative interval).
What is a DOI Number?

Some links on this page may take you to non-federal websites. Their policies may differ from this site.

Melody: meta-analysis of microbiome association studies for discovering generalizable microbial signatures

https://doi.org/10.1186/s13059-025-03721-4

Wei, Zhoujingpeng; Chen, Guanhua; Tang, Zheng-Zheng (August 2025, Genome Biology)

Abstract Standard protocols for meta-analysis of association studies are inadequate for microbiome data due to their complex compositional structure, leading to inaccurate and unstable microbial signature selection. To address this issue, we introduce Melody, a framework that generates, harmonizes, and combines study-specific summary association statistics to powerfully and robustly identify microbial signatures in meta-analysis. Comprehensive and realistic simulations demonstrate that Melody substantially outperforms existing approaches in prioritizing true signatures. In the meta-analyses of five studies on colorectal cancer and eight studies on the gut metabolome, we showcase the superior stability, reliability, and predictive performance of Melody-identified signatures.
more » « less
Multidimensional scaling improves distance-based clustering for microbiome data

https://doi.org/10.1093/bioinformatics/btaf042

Chen, Guanhua; Wang, Xinyue; Sun, Qiang; Tang, Zheng-Zheng (February 2025, Bioinformatics)
Birol, Inanc (Ed.)
Abstract Motivation:Clustering patients into subgroups based on their microbial compositions can greatly enhance our understanding of the role of microbes in human health and disease etiology. Distance-based clustering methods, such as partitioning around medoids (PAM), are popular due to their computational efficiency and absence of distributional assumptions. However, the performance of these methods can be suboptimal when true cluster memberships are driven by differences in the abundance of only a few microbes, a situation known as the sparse signal scenario. Results:We demonstrate that classical multidimensional scaling (MDS), a widely used dimensionality reduction technique, effectively denoises microbiome data and enhances the clustering performance of distance-based methods. We propose a two-step procedure that first applies MDS to project high-dimensional microbiome data into a low-dimensional space, followed by distance-based clustering using the low-dimensional data. Our extensive simulations demonstrate that our procedure offers superior performance compared to directly conducting distance-based clustering under the sparse signal scenario. The advantage of our procedure is further showcased in several real data applications. Availability and implementation:The R package MDSMClust is available at https://github.com/wxy929/MDS-project.
more » « less
Free, publicly-accessible full text available February 1, 2026
Fair prediction of 2-year stroke risk in patients with atrial fibrillation

https://doi.org/10.1093/jamia/ocae170

Gao, Jifan; Mar, Philip; Tang, Zheng-Zheng; Chen, Guanhua (July 2024, Journal of the American Medical Informatics Association)

Abstract ObjectiveThis study aims to develop machine learning models that provide both accurate and equitable predictions of 2-year stroke risk for patients with atrial fibrillation across diverse racial groups. Materials and MethodsOur study utilized structured electronic health records (EHR) data from the All of Us Research Program. Machine learning models (LightGBM) were utilized to capture the relations between stroke risks and the predictors used by the widely recognized CHADS2 and CHA2DS2-VASc scores. We mitigated the racial disparity by creating a representative tuning set, customizing tuning criteria, and setting binary thresholds separately for subgroups. We constructed a hold-out test set that not only supports temporal validation but also includes a larger proportion of Black/African Americans for fairness validation. ResultsCompared to the original CHADS2 and CHA2DS2-VASc scores, significant improvements were achieved by modeling their predictors using machine learning models (Area Under the Receiver Operating Characteristic curve from near 0.70 to above 0.80). Furthermore, applying our disparity mitigation strategies can effectively enhance model fairness compared to the conventional cross-validation approach. DiscussionModeling CHADS2 and CHA2DS2-VASc risk factors with LightGBM and our disparity mitigation strategies achieved decent discriminative performance and excellent fairness performance. In addition, this approach can provide a complete interpretation of each predictor. These highlight its potential utility in clinical practice. ConclusionsOur research presents a practical example of addressing clinical challenges through the All of Us Research Program data. The disparity mitigation framework we proposed is adaptable across various models and data modalities, demonstrating broad potential in clinical informatics.
more » « less
PhyloMed: a phylogeny-based test of mediation effect in microbiome

https://doi.org/10.1186/s13059-023-02902-3

Hong, Qilin; Chen, Guanhua; Tang, Zheng-Zheng (December 2023, Genome Biology)

Abstract Microbiome data from sequencing experiments contain the relative abundance of a large number of microbial taxa with their evolutionary relationships represented by a phylogenetic tree. The compositional and high-dimensional nature of the microbiome mediator challenges the validity of standard mediation analyses. We propose a phylogeny-based mediation analysis method called PhyloMed to address this challenge. Unlike existing methods that directly identify individual mediating taxa, PhyloMed discovers mediation signals by analyzing subcompositions defined on the phylogenic tree. PhyloMed produces well-calibrated mediation test p -values and yields substantially higher discovery power than existing methods.
more » « less
Full Text Available
A reluctant additive model framework for interpretable nonlinear individualized treatment rules

https://doi.org/10.1214/23-AOAS1767

Maronge, Jacob M; Huling, Jared D; Chen, Guanhua (December 2023, The Annals of Applied Statistics)

Individualized treatment rules (ITRs) for treatment recommendation is an important topic for precision medicine as not all beneficial treatments work well for all individuals. Interpretability is a desirable property of ITRs, as it helps practitioners make sense of treatment decisions, yet there is a need for ITRs to be flexible to effectively model complex biomedical data for treatment decision making. Many ITR approaches either focus on linear ITRs, which may perform poorly when true optimal ITRs are nonlinear, or blackbox nonlinear ITRs, which may be hard to interpret and can be overly complex. This dilemma indicates a tension between interpretability and accuracy of treatment decisions. Here we propose an additive model-based nonlinear ITR learning method that balances interpretability and flexibility of the ITR. Our approach aims to strike this balance by allowing both linear and nonlinear terms of the covariates in the final ITR. Our approach is parsimonious in that the nonlinear term is included in the final ITR only when it substantially improves the ITR performance. To prevent overfitting, we combine crossfitting and a specialized information criterion for model selection. Through extensive simulations we show that our methods are data-adaptive to the degree of nonlinearity and can favorably balance ITR interpretability and flexibility. We further demonstrate the robust performance of our methods with an application to a cancer drug sensitive study.
more » « less
Full Text Available
High-throughput screening of dual atom catalysts for oxygen reduction and evolution reactions and rechargeable zinc-air battery

https://doi.org/10.1016/j.nanoen.2024.109634

Tamtaji, Mohsen; Kim, Min Gyu; Li, Zhimin; Cai, Songhua; WANG, Jun; Galligan, Patrick Ryan; Hung, Faan-Fung; Guo, Hui; Chen, Shuguang; Luo, Zhengtang; et al (July 2024, Nano Energy)

Full Text Available
A High‐Entropy Single‐Atom Catalyst Toward Oxygen Reduction Reaction in Acidic and Alkaline Conditions

https://doi.org/10.1002/advs.202309883

Tamtaji, Mohsen; Kim, Min Gyu; WANG, Jun; Galligan, Patrick Ryan; Zhu, Haoyu; Hung, Faan‐Fung; Xu, Zhihang; Zhu, Ye; Luo, Zhengtang; Goddard, William A; et al (April 2024, Advanced Science)

Abstract The design of high‐entropy single‐atom catalysts (HESAC) with 5.2 times higher entropy compared to single‐atom catalysts (SAC) is proposed, by using four different metals (FeCoNiRu‐HESAC) for oxygen reduction reaction (ORR). Fe active sites with intermetallic distances of 6.1 Å exhibit a low ORR overpotential of 0.44 V, which originates from weakening the adsorption of OH intermediates. Based on density functional theory (DFT) findings, the FeCoNiRu‐HESAC with a nitrogen‐doped sample were synthesized. The atomic structures are confirmed with X‐ray photoelectron spectroscopy (XPS), X‐ray absorption (XAS), and scanning transmission electron microscopy (STEM). The predicted high catalytic activity is experimentally verified, finding that FeCoNiRu‐HESAC has overpotentials of 0.41 and 0.37 V with Tafel slopes of 101 and 210 mVdec⁻¹at the current density of 1 mA cm⁻²and the kinetic current densities of 8.2 and 5.3 mA cm⁻², respectively, in acidic and alkaline electrolytes. These results are comparable with Pt/C. The FeCoNiRu‐HESAC is used for Zinc–air battery applications with an open circuit potential of 1.39 V and power density of 0.16 W cm⁻². Therefore, a strategy guided by DFT is provided for the rational design of HESAC which can be replaced with high‐cost Pt catalysts toward ORR and beyond.
more » « less
Full Text Available
Causal Inference Methods for Combining Randomized Trials and Observational Studies: A Review

https://doi.org/10.1214/23-STS889

Colnet, Bénédicte; Mayer, Imke; Chen, Guanhua; Dieng, Awa; Li, Ruohong; Varoquaux, Gaël; Vert, Jean-Philippe; Josse, Julie; Yang, Shu (February 2024, Statistical Science)

Full Text Available
Independence Weights for Causal Inference with Continuous Treatments

https://doi.org/10.1080/01621459.2023.2213485

Huling, Jared D.; Greifer, Noah; Chen, Guanhua (July 2023, Journal of the American Statistical Association)

Full Text Available
Carrier-Envelope-Phase Modulated Currents in Scanning Tunneling Microscopy

https://doi.org/10.1021/acs.nanolett.1c01900

Hu, Ziyang; Kwok, YanHo; Chen, GuanHua; Mukamel, Shaul (August 2021, Nano Letters)

Full Text Available

« Prev Next »

Search for: All records